Papers with multilingual learning

11 papers
Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER (P18-2)

Copied to clipboard

Challenge: Existing approaches to improve NER performance add training data from one or more assisting languages to the primary language.
Approach: They propose a metric based on symmetric KL divergence to filter out highly divergent training instances in the assisting language.
Outcome: The proposed method improves NER performance in many languages, including those with limited training data.
75 Languages, 1 Model: Parsing Universal Dependencies Universally (D19-1)

Copied to clipboard

Challenge: UDify is a multilingual multi-task model that can predict universal part-of-speech, morphological features, lemmas, and dependency trees.
Approach: They evaluate UDify, a multilingual multi-task model capable of predicting universal part-of-speech, morphological features, lemmas, and dependency trees simultaneously for all 124 Universal Dependencies treebanks across 75 languages.
Outcome: The proposed model can predict universal part-of-speech, morphological features, lemmas, and dependency trees for all 124 treebanks across 75 languages.
X-SRL: A Parallel Cross-Lingual Semantic Role Labeling Dataset (2020.emnlp-main)

Copied to clipboard

Challenge: Existing multilingual SRL datasets contain disparate annotation styles or come from different domains, hampering generalization in multilingual learning.
Approach: They propose to automatically construct an SRL corpus that is parallel in four languages with unified predicate and role annotations that are fully comparable across languages.
Outcome: The proposed method improves performance for English SRL in weaker languages.
Offensive language detection in Hebrew: can other languages help? (2022.lrec-1)

Copied to clipboard

Challenge: Various approaches for offensive language detection have been applied for this task . contamination of social networks with offensive content is a new reality affecting almost all of us .
Approach: They propose to use multiple supervised models and text representations to detect offensive language in three languages, including two Semitic languages.
Outcome: The proposed model can detect offensive content in two Semitic languages, including Hebrew and Arabic, and it is able to perform cross-lingual and multilingual learning.
A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking (2020.lrec-1)

Copied to clipboard

Challenge: Existing multimodal corpora lack the ability to be used in multilingual or non-English scenarios.
Approach: They extend a Flickr30k Entities image-caption dataset with Japanese translations to provide a multilingual corpus.
Outcome: The proposed dataset is the first multilingual image-caption dataset with Japanese translations.
Hyper-X: A Unified Hypernetwork for Multi-Task Multilingual Transfer (2022.emnlp-main)

Copied to clipboard

Challenge: Existing multilingual models cannot fully leverage training data when it is available in different task-language combinations.
Approach: They propose a single hypernetwork that unifies multi-task and multilingual learning with efficient adaptation.
Outcome: The proposed model achieves the best or competitive gain when a mixture of multiple resources is available while being significantly more efficient than existing models.
Don’t Go Far Off: An Empirical Study on Neural Poetry Translation (2021.emnlp-main)

Copied to clipboard

Challenge: despite improvements in machine translation quality, automatic poetry translation remains a challenging problem . et al., a study of automatic poetry translators shows that multilingual fine-tuning on poetic data outperforms bilingual fine-timing on non-poetic text .
Approach: They propose to use poetic parallel corpora for 6 languages to study poetry translation . they find that multilingual fine-tuning on poetic data outperforms bilingual fine-uning .
Outcome: The proposed model outperforms bilingual and multilingual models on poetic data . the proposed model is based on a parallel dataset of poetry translations for several languages .
Polyglot Prompt: Multilingual Multitask Prompt Training (2022.emnlp-main)

Copied to clipboard

Challenge: a monolithic framework for multilingual learning can be used without any task/language-specific module.
Approach: They propose a framework to exploit prompting methods for learning a unified semantic space for different languages and tasks with multilingual prompt engineering.
Outcome: The proposed framework can learn tasks from different languages in a monolithic framework without any task/language-specific module.
Multilingual Transfer Learning for Children Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in automatic speech recognition (ASR) systems have been criticized for high acoustic variability and limited amount of available training data.
Approach: They propose a two-step training strategy that uses multilingual learning followed by language-specific transfer learning to generalize children's speech.
Outcome: The proposed training strategy outperforms single language training and multilingual and transfer learning alone in English.
Typology Guided Multilingual Position Representations: Case on Dependency Parsing (2023.findings-acl)

Copied to clipboard

Challenge: Recent multilingual models benefit from strong unified semantic representation models, but conflicting linguistic regularities may break the effectiveness of word position features in multilingual learning.
Approach: They propose to combine prior knowledge from typology features and existing position vectors to create a position generation network which combines prior knowledge of a language's position space and typological characterization.
Outcome: The proposed model can achieve the best multilingual parsing results by combining prior knowledge from typology features and existing position vectors.
ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have led to significant breakthroughs in the field.
Approach: They evaluate ChatGPT over multiple tasks with diverse languages and large datasets to provide more comprehensive information for multilingual NLP applications.
Outcome: The proposed model can process and generate texts for multiple languages due to its multilingual training data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations